An Adaptive Binarization Technique for Low Quality Historical Documents

نویسندگان

  • Basilios Gatos
  • Ioannis Pratikakis
  • Stavros J. Perantonis
چکیده

Historical document collections are a valuable resource for human history. This paper proposes a novel digital image binarization scheme for low quality historical documents allowing further content exploitation in an efficient way. The proposed scheme consists of five distinct steps: a pre-processing procedure using a low-pass Wiener filter, a rough estimation of foreground regions using Niblack’s approach, a background surface calculation by interpolating neighboring background intensities, a thresholding by combining the calculated background surface with the original image and finally a post-processing step in order to improve the quality of text regions and preserve stroke connectivity. The proposed methodology works with great success even in cases of historical manuscripts with poor quality, shadows, nonuniform illumination, low contrast, large signal-dependent noise, smear and strain. After testing the proposed method on numerous low quality historical manuscripts, it has turned out that our methodology performs better compared to current state-of-the-art adaptive thresholding techniques.

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

An Enhancement of Images Using Recursive Adaptive Gamma Correction

The “Adaptive Approach for Historical or Degraded Document Binarization” is that in which Libraries and Museums obtain in large gathering of ancient historical documents printed or handwritten in native languages. Typically, only a small group of people are allowed access to such collection, as the preservation of the material is of great concern. In recent years, libraries have begun to digiti...

متن کامل

Restoration of Degraded Historical Document Image: An Adaptive Multilayer-Information Binarization Technique

Binary image is the essential format for document image processing, and the operation of the subsequent steps depends on the quality of the binarization process. The objective of this research is to propose a new binarization method based on adaptive multilayer-information for restoration of degraded historical document images. This paper focuses on degraded Thai historical document images, whi...

متن کامل

A Proposed Binarization Technique on Hand written document

Abstract: Binarization is performed in the preprocessing stage for document inspection. Binarization of degraded document images improve the result from poor quality of the paper, the printing process, ink blot and fading document and remove noise from examine. In recent years, libraries have begun to digitize historical document that are of interest to a wide range of people, with the goal of ...

متن کامل

Hybrid Binariztion Technique for Historical Manuscripts

This paper presents a new hybrid approach for the binarization and enhancement of Historical Manuscript. This paper deals with degradations which occur due to shadows, non-uniform illumination, low contrast and strain. We follow two distinct method of Binarization with a pre-processing procedure using a adaptive Wiener filter, a rough estimation of foreground regions and a background surface ca...

متن کامل

Locating Text in Historical Collection Manuscripts

It is common that documents belonging to historical collections are poorly preserved and are prone to degradation processes. The aim of this work is to leverage state-of-the-art techniques in digital image binarization and text identification for digitized documents allowing further content exploitation in an efficient way. A novel methodology is proposed that leads to preservation of meaningfu...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

برای دانلود متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

عنوان ژورنال:

دوره   شماره 

صفحات  -

تاریخ انتشار 2004